Back

Acta Crystallographica Section D Structural Biology

International Union of Crystallography (IUCr)

Preprints posted in the last 7 days, ranked by how well they match Acta Crystallographica Section D Structural Biology's content profile, based on 59 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.

1
Cryo-EM structure of CYP2C9 reveals a dimer-of-trimers assembly

Tanino, H.; Tsujino, H.; Nakao, T.; Oie, C.; Makino, F.; Miyata, T.; Kasai, K.; Namba, K.; Inoue, T.

2026-08-31 biophysics 10.64898/2026.08.30.747968 medRxiv
Top 0.2%
2.8%
Show abstract

Human cytochrome P450 2C9 (CYP2C9) is a hepatic microsomal enzyme involved in the oxidative metabolism of clinically important drugs, but the structural organization of its oligomeric assemblies outside crystallographic packing environments remains poorly understood. Here, we report the cryo-EM structure of human CYP2C9 determined under aqueous, membrane-free conditions at 3.31 Angstrom resolution. The structure reveals a C2-symmetric hexameric assembly organized as a dimer of trimers. Individual protomers retain the conserved P450 fold and heme-binding architecture observed in previously reported crystal structures, indicating that assembly formation does not substantially perturb the catalytic core. The hexamer is stabilized by defined intra-trimer interfaces involving the N-terminal region and residues around Trp212 and Phe482, together with inter-trimer interfaces involving Leu71 and the 220-227 loop. These interfaces are distinct from the crystal packing contacts observed in CYP2C9 crystal structures, demonstrating that the assembly is not a simple recapitulation of crystallographic packing. Notably, the inter-trimer interface is located near the FG-loop-containing surface previously implicated in membrane association. This suggests that the observed hexamer may represent a membrane-free association of two trimers through membrane-related surfaces, whereas the trimeric arrangement itself may be compatible with membrane-associated organization. The structure therefore provides a framework for investigating how trimer formation, membrane interaction and local conformational changes in the FG-loop region may influence CYP2C9 function.

2
The Role Of Liquid Crystal Ordering In The Structural Organization Of DNA In Bacteria.

Krupyanskii, Y. F.; Kovalenko, V.; Loiko, N.; Generalova, A.; Tereshkin, E.; Tereshkina, K.; Sokolova, O.; Peters, G.

2026-09-01 biophysics 10.64898/2026.08.31.748243 medRxiv
Top 0.2%
2.7%
Show abstract

This paper presents and critically reviews the results of original and some literature based experimental studies conducted by the authors last years on the structural organization of DNA in dormant (starvation stress), anabiotic dormant (4 HR treatment) E. coli cells, as well as the K12 {Delta}dps strain, which lacks the Dps protein (Dps null E. coli). The experimental data includes small-angle synchrotron radiation diffraction (SAXS) and transmission electron microscopy (TEM) data. Synchrotron radiation diffraction experiments on K12{Delta}dps cells allowed us to conclude that peaks at 44.3, 22.1, and 14.8 angstrom resolutions are associated exclusively with ordered DNA organization. Peaks at 44.3, 22.1, and 14.8 angstrom resolutions are also observed for samples of dormant (starvation stress) cells and anabiotically dormant cells. Therefore, this ordered DNA organization also applies to samples of dormant and anabiotically dormant cells. A model is proposed that considers the ordered DNA organization in the cell as a cholesteric liquid crystal. The powder diffraction pattern calculated based on this model is compared with experimental small angle X ray scattering (SAXS) data obtained on Dps-null cell samples. The model completely reproduces the key features of the experimental diffraction pattern from Dps-null cell samples. Accordingly, the cholesteric liquid crystal model corresponds to DNA packaging in dormant and anabiotically dormant cells. Cholesteric liquid crystal ordering should be further considered in all models of cellular DNA packaging. To address the question of which structural organization of DNA predominates in the cell: the cholesteric liquid crystal or nanocrystalline or whether they coexist and fully manifest themselves under different external conditions, it is necessary to utilize the latest methodological advances in structural analysis.

3
Pre-FIB Layer-Mapping Cryo Tomography (PLCT) for Depth-Resolved in Situ Structural Analysis of Multilayered Tissues

Wang, F.; Lin, X.; Rao, B.; Lai, X.; Yu, L.; Sun, F.; Qu, J.; Zhang, J.

2026-08-30 neuroscience 10.64898/2026.08.25.746966 medRxiv
Top 0.7%
0.5%
Show abstract

Cryo-electron tomography (cryo-ET) enables near-native visualization of subcellular architectures, yet applying it to moderately thick, multilayered tissues such as the retina is hampered by inadequate vitrification and inaccurate depth-targeting. Here, we developed PLCT, an integrated approach combining modified high-pressure freezing, cryo-ultramicrotome trimming, and plasma-based cryo-FIB milling to overcome these barriers. PLCT reliably vitrified <100 m retinal strips with minimal ice artifacts, navigates precisely to the outer plexiform layer using morphological landmarks, and produces high-quality lamellae suitable for high-resolution cryo-ET. Subtomogram averaging (STA) analysis identified microtubules at 16.33 [A] within retinal horizontal cell processes. Importantly, STA also resolved a 10-nm-diameter filamentous structure at 24.81 [A] in the same processes, featuring six peripheral strands surrounding an elongated central density with continuous intervening cavities, an architecture consistent with intermediate filaments. Together with its native localization and immunoreactivity, these features collectively identify the filaments as neurofilaments. Separately, 3D reconstruction of synaptic ribbons uncovered a previously unrecognized "mahjong tile"-like fine ultrastructure. These results demonstrate that PLCT-produced lamellae are of sufficient quality to support structural analysis in native tissue. Although demonstrated on retinal photoreceptor synapses as a proof-of-principle, PLCT is inherently generalizable, with its depth-navigation and vitrification strategies directly applicable to any multilayered tissues. This work establishes PLCT as a robust, reproducible platform for depth-resolved in situ cryo-ET of multilayered tissues.

4
Cryo-EM Structure of a Triazole alpha-Conotoxin GI Mimetic Bound to the Muscle-Type Nicotinic Acetylcholine Receptor

Shepperson, O.; Capper, M.; Holdship, C.; Melling, O.; Wade, N.; Malone, M.; Arnott, K.; Morgan, D.; Piggot, T.; Morcom, T.; Connah, J.; Windeln, L.; Timperley, C.; Frey, J.; Green, C.; Koehnke, J.; Essex, J.; Jamieson, A.

2026-09-01 biochemistry 10.64898/2026.08.31.748223 medRxiv
Top 0.7%
0.5%
Show abstract

Disulfide-rich peptides possess exceptional potency and selectivity but are often limited by the instability and synthetic challenges associated with native disulfide bonds. Here, we report the design, synthesis, pharmacological evaluation, and structural characterisation of triazole-based peptidomimetics of the -GI conotoxin, a selective antagonist of the muscle-type nicotinic acetylcholine receptor (nAChR). A series of 1,4- and 1,5-disubstituted triazole analogues were prepared entirely on resin using CuAAC and RuAAC chemistry to replace the native Cys3/13 disulfide bridge. Functional evaluation against human muscle nAChRs revealed that 1,5-triazole analogues retained low-nanomolar potency, with the lead mimetic exhibiting activity comparable to native -GI. Cryo-electron microscopy of the lead compound bound to the muscle-type nAChR provided the first structure of a disulfide-isostere peptidomimetic in complex with a membrane receptor. The structure demonstrates that the 1,5-triazole reproduces the native peptide fold with high fidelity while contributing receptor-facing interactions not available to the native disulfide bridge. Molecular dynamics simulations further revealed conserved hydration networks and similar conformational sampling between the native peptide and lead mimetic. Together, these findings establish triazoles as effective disulfide surrogates and provide a structural framework for the rational design of stabilised conotoxin therapeutics.

5
Structural and biochemical analysis of the IBV nsp15 endoribonuclease reveals the necessity of peripheral site residues for activity

Tully, E. S.; Kirchdoerfer, R. N.

2026-08-31 biochemistry 10.64898/2026.08.28.747898 medRxiv
Top 0.7%
0.5%
Show abstract

Infectious bronchitis virus (IBV) is a member of the Gammacoronavirus genus responsible for respiratory illness and weakened eggshells in infected chickens, adversely impacting the poultry industry. Escaping innate immune detection during infection is crucial for coronavirus proliferation in the host. The production of double-stranded RNA during coronavirus replication triggers innate immune sensors to create an antiviral state within infected cells. To counter this response, coronaviruses employ nonstructural protein 15 (nsp15) endoribonuclease to degrade double-stranded RNA. Here, we use cryo-electron microscopy and biochemistry to characterize IBV nsp15 interactions with RNA. While the overall structure and active site of IBV nsp15 strongly resemble previous studies of nsp15 from other coronaviral genera, we note that double-stranded RNA contacts several non-conserved residues peripheral to the enzyme active site. Our data show that these residue positions can have strong impacts on RNA cleavage suggesting unique solutions for RNA engagement across coronavirus species. We also demonstrate a preference for IBV nsp15 to cleave double-stranded RNA over single-stranded RNA and observe nsp15 hexamers with two double-stranded RNAs bound simultaneously. This study reinforces the need to study diverse coronavirus species to identify distinct viral enzyme characteristics.

6
Structural basis for catalytic and inhibitory divergence between archaeal and bacterial ammonia monooxygenases

Yang, X.; Mao, T.-Q.; He, Z.-C.; Chen, Y.; Zhao, G.; Jin, P.; Li, S.; Dong, H.-P.; Peng, W.; Zhang, C.; Li, Z.

2026-09-01 molecular biology 10.64898/2026.08.31.748207 medRxiv
Top 0.8%
0.4%
Show abstract

Ammonia oxidation initiates nitrification and is closely linked to microbial N2O production. Ammonia monooxygenase (AMO) catalyzes the first and rate-limiting step of nitrification and is widespread across evolutionarily distinct ammonia-oxidizing archaea (AOA) and bacteria (AOB). The ocean is the largest biome for AOA and AOB, which have distinct ecological niches and markedly different sensitivities to nitrification inhibitors. However, the lack of archaeal AMO structures and inhibitor-bound AMO complexes has hindered mechanistic understanding of the architectural, catalytic, and inhibitory divergence between these two enzyme systems. Here, we report high-resolution cryo-electron microscopy (cryo-EM) structures of marine archaeal AMO captured in active and inactivated states within its native membrane environment, together with inhibitor-bound structures of estuarine bacterial AMO. Archaeal AMO forms an unexpected cup-shaped homotrimer composed of eight subunits per protomer and exhibits substantial architectural divergence from bacterial AMO. Integrated structural, biochemical, kinetic, and computational analyses reveal distinct periplasmic architectures, copper-center organization, and hydrophobic channels between archaeal and bacterial AMOs for ammonium acquisition, catalysis and inhibitor response. These findings provide a structural and mechanistic framework for understanding how archaeal and bacterial AMOs have diverged to distinct ammonia-oxidizing strategies and inhibitor susceptibilities across environmentally important ammonia oxidizers.

7
A transition state-like acylenzyme conformation distinguishes carbapenemase activity in class A β-lactamases

Beer, M.; Spencer, J.; Mulholland, A. J.

2026-09-01 biochemistry 10.64898/2026.08.31.748333 medRxiv
Top 0.8%
0.4%
Show abstract

Carbapenems are the most potent {beta}-lactams, key antibiotics for healthcare-associated infections by Gram-negative bacteria and evade hydrolysis by most {beta}-lactamases, but are increasingly threatened by emergence of enzymes exhibiting hydrolytic activity towards them. Of the four recognised {beta}-lactamase subclasses, class A (active-site serine enzymes that hydrolyse {beta}-lactams via a covalent acylenzyme intermediate) is the most widely disseminated and, while the majority of such enzymes react with carbapenems to form long-lasting acylenzyme complexes, several possess carbapenem-hydrolyzing activity (carbapenemases). Here, we investigate the basis for these differences in a panel of class A {beta}-lactamases using molecular dynamics (MD) simulations of the respective acylenzyme complexes and tetrahedral intermediates (TI). The simulations reveal multiple features associated with catalytic activity across the spectrum of enzymes tested, including more extensive interactions of the carbapenem acylenzyme carbonyl and generally increased lifetimes of active site water molecules positioned for deacylation. Analysis of the dynamic trajectories shows carbapenemases to have reduced root mean-squared fluctuation (RMSF) differences between the acylenzyme and TI, that are not limited to the active site, indicating that the acylenzyme complex is pre-organised for reaction in carbapenemases but not in carbapenem-inhibited enzymes. Similarly, Principal Component Analysis (PCA) of acylenzyme and TI dynamics shows greater overlap between the two states in carbapenemases, providing further evidence for acylenzyme pre-organisation. Such simulations may represent an effective computational assay able to identify enzymes with carbapenemase activity at relatively modest computational cost.

8
Polarized neutrons for the study of individual and collective fast dynamics in proteins

Nidriche, A.; Ollivier, J.; Stewart, R.; Peters, J.

2026-09-01 biophysics 10.64898/2026.08.30.748099 medRxiv
Top 1%
0.3%
Show abstract

Neutron scattering is a powerful technique to investigate atomic structures and molecular dynamics of proteins at the nano-scale. When it comes to dynamics, incoherent and coherent scattering respectively provide information on the single and collective dynamics of nuclei. In proteins, hydrogen has the highest incoherent cross-section, and it is common practice to overlook the contribution of coherent terms stemming from all nuclei. However, the fast collective dynamics of heavier nuclei could also be studied if coherent scattering and incoherent scattering were experimentally separated. The recent advent of polarized neutron spectroscopy with sufficient flux and energy resolution has made it possible, and opens new perspectives to investigate the relative importance of coherent scattering and the information it provides on biological samples. The present study reports on the use of polarized quasi-elastic neutron scattering (QENS) and the application of a minimalistic model adapted to both individual and collective dynamics. Using a perdeuterated green fluorescent protein as a model globular protein, the study provides an interpretation of the dynamical parameters obtained with QENS, and a comparative study of the Elastic Coherent and Incoherent Scattering Factor. Based on both experiments and calculations, we discuss the relative importance of distinct and self components of coherent scattering, which is often wrongly assumed to be representative of collective dynamics only. The results highlight the current impediments rendering complicated a straightforward analysis of fast collective dynamics in hydrated protein samples.

9
Identification and structural basis of a Chloroflexus protein with homology to Bacillus quorum sensing-related prenyltransferase

Matsui, T.; Inoue, S.; Yanagimoto, S.; Kaneko, A.; Tago, R.; Suto, A.; Odagi, M.; Kodera, Y.; Morita, H.; Abe, I.; Okada, M.

2026-08-31 biochemistry 10.64898/2026.08.29.745113 medRxiv
Top 1%
0.3%
Show abstract

Quorum sensing in Gram-positive bacteria commonly relies on posttranslationally modified peptide pheromones. In Bacillus subtilis, the prenyltransferase ComQ catalyzes tryptophan prenylation of the quorum-sensing peptide ComX, but the structural basis of this unique peptide modification has remained unclear. Here we identified a previously uncharacterized ComQ homolog, StheQ, and its cognate peptide substrate, StheX, from Sphaerobacter thermophilus and investigated their structural and functional relationship. Liquid chromatography-tandem mass spectrometry (LC-MS/MS) analysis demonstrated that StheQ catalyzes prenylation of the tryptophan residue located second from the C-terminus of StheX. Crystal structures of apo StheQ and its complexes with a farnesyl pyrophosphate analog revealed that StheQ adopts the all--helical fold of the trans-isoprenyl diphosphate synthase (IPPS) superfamily while possessing an active-site architecture adapted for peptide-based indole prenylation. The structures identified a single Mg2+-binding site associated with the first aspartic acid-rich motif and showed no evidence for metal coordination at the pseudo-second aspartic acid-rich motif. Site-directed mutagenesis, complex formation assays, and docking analyses identified a peptide-binding pocket adjacent to the active site and suggested that N215 contributes to productive positioning of the acceptor tryptophan. These findings establish the structural basis for peptide prenylation by a ComQ-family enzyme, providing insight into the evolution of peptide-based indole prenylation within the IPPS superfamily, and support the view that ComQ-family enzymes constitute a distinct functional branch specialized for peptide modification.

10
The first OpenBind release: An open experimental structure-affinity dataset and benchmark for structure-based AI

Nelen, J.; Khan, O.; Adams, E.; Aschenbrenner, J. C.; Thompson, W.; Ebrahim, A.; Capkin, E.; Vallee, C.; OpenBind, ; Shotton, E. J.; Griffen, E. J.; Chodera, J. D.; Deane, C. M.; von Delft, F.; AlQuraishi, M.; Imrie, F.

2026-09-01 bioinformatics 10.64898/2026.08.27.747600 medRxiv
Top 1%
0.3%
Show abstract

High-quality experimental datasets that link protein-ligand structures with binding affinity data are essential for developing and evaluating structure-based machine learning methods. To help address this need, we established OpenBind as an open-science initiative to generate large-scale experimental datasets for structure-based AI and molecular discovery. Here, we describe the first public OpenBind release, which, to the best of our knowledge, is the largest public single-target experimental structure-affinity dataset. The dataset focuses on enteroviral 2A protease, comprising 925 crystallographic binding events from 699 compounds and associated affinity measurements for 601 compounds. It combines structures from an initial fragment screen and follow-on molecules, together with affinity data, linking experimentally determined protein-ligand binding modes to biophysical measurements within a coherent antiviral discovery campaign. We used this dataset to evaluate protein-ligand structure prediction, binding-affinity prediction, and virtual screening using representative structure-based methods, including docking and cofolding. This exposed several challenges that are central to practical structure-based modelling: docking performance depends strongly on binding-pocket conformation, poses are difficult to rank, and structure-based affinity prediction remains challenging. Fine-tuning OpenFold3-p2 on the fragment-screen structures substantially improved pose prediction and virtual screening for related follow-on compounds, demonstrating how early-stage experimental structures can support target-specific model adaptation.

11
From Prompt to Provenance: BloClaw, a Capability-Gated AI4S Workstation for Auditable Computational Biology

qin, y.; Pang, J.; Zhang, X.

2026-09-01 bioinformatics 10.64898/2026.08.26.747436 medRxiv
Top 1%
0.3%
Show abstract

Scientific agents can produce plausible answers while remaining unable to establish whether the computation behind an answer is executable, recoverable, or reproducible. We present BloClaw, an AI4S workstation built around a simple principle: a scientific agent should know what it can do, show how it did it, and state what remains unvalidated. Each capability declares an execution state, input constraints, dependencies, expected outputs, and scientific limitations. Natural-language requests are translated into structured tasks, validated against this registry, executed through scientific tools, and recorded in a provenance-aware Living Lab Notebook. The system is designed to detect invalid inputs, failed tool calls, missing dependencies, and remote timeouts, and to route them to repair, retry, or escalation. The implemented and tested scope comprises RDKit-based molecular property and rule screening, protein structure analysis, docking-pose inspection, 3D visualization, and structured reporting. We demonstrate the workflow on a PubChem-retrieved osimertinib structure and a supplied 6LU7 docking artifact: the former yields deterministic descriptors (molecular weight 499.619 Da, cLogP 4.5098, TPSA 87.55 A^2), while the latter contains 2,387 protein ATOM records, 309 residues, and nine pose records. These examples are workflow demonstrations, not efficacy or affinity studies. Beyond retrospective prediction, the manuscript specifies a prior-minimized constructive mode in which a desired function is compiled into explicit physical, chemical, and systems constraints, candidate mechanisms are simulated, and observations are reintroduced for calibration and falsification; this is a proposed extension rather than a result of the present case studies. We describe an evaluation protocol that compares BloClaw with a standard single-agent workflow and fixed-script execution using task completion, scientific correctness, recovery success, provenance completeness, reproducibility, human review time, latency, and cost. This manuscript reports the system design, verified capability boundary, deterministic software artifacts, and a reproducible evaluation protocol; it does not claim benchmark improvements before those experiments are run. BloClaw is an execution and accountability layer for AI-assisted research, complementing expert review and experimental validation rather than replacing them.

12
The interaction between NC(p7)1-55 and p6 may regulate interactions with nucleic acids during assembly through modulation of Gag folding.

LARUE, V.; Nonin-Lecomte, S.

2026-09-01 biophysics 10.64898/2026.08.28.747767 medRxiv
Top 1%
0.3%
Show abstract

We present the solution structures of HIV-1 proteins NC(p7)1-55 corresponding to the full-length NC(p7) and mature p6. The studies were carried in water and, to mimic the membrane, in micellar DPC (Dodecylphosphocholine) conditions. Our results unravel for the first time the structure adopted by the N-terminal amino acids of the free NC(p7)1-55, with the formation of a small helix spanning residues F6 to R10. Our NMR and Fluorescence Anisotropy data disclose an interaction between NC(p7)1-55 and p6 both in water and DPC, with respective Kd of 2.5mM and 370 mM at 23{degrees}C. The interaction is thus strengthened in lipidic conditions. Protein p6 stabilizes the N-terminus of NC(p7)1-55 while increasing at the same time the dynamic of the first zinc finger. Although the entire p6 sequence is involved in the interaction, we show that its C-terminal region is particularly sensitive to the presence of NC(p7)1-55, with a propensity of forming a a helix ranging from amino acids S111 to F116. This study brings experimental evidence of a direct protein-protein interaction between p6 and the N-terminal region of NC(p7)1-55. We further show that such interaction is readily accommodated within the NC(p15) framework and hypothesize that it may facilitate the selective assembly of assembly of the viral genomic RNA (gRNA) in the cell.

13
Structural characterization the LlaI anti-phage defense system reveals insights into the evolution of nucleotide specificity and the organization of DNA binding in McrBC restriction complexes

Bui, A. Q.; Hosford, C. J.; Niu, Y.; Santiago, E.; Moraga, D.; Wagner, M. M.; Chappie, J. S.

2026-09-01 biochemistry 10.64898/2026.08.31.748284 medRxiv
Top 1%
0.2%
Show abstract

Canonical McrBC enzymes are nucleotide-powered, motor-driven endonucleases that bind and cleave modified bacteriophage DNA. Non-canonical McrBC homologs like LlaI and BsuMI are distinguished by a unique three-gene organization and the ability to target DNA site-specifically. Here, we report the atomic-resolution crystal structures of the DNA-binding module LlaI.R1 and AAA+ motor LlaI.R2 from the Lactococcus lactis LlaI anti-phage defense system. The crystallized LlaI.R2 hexamer traps two distinct active site conformations that correlate to different states of the nucleotide hydrolysis cycle and reveal that the organization of the critical catalytic machinery present in canonical McrB homologs is also conserved in non-canonical R2 proteins. Although canonical McrB homologs are strictly GTP-specific, we find that the R2 proteins from LlaI and BsuMI do not discriminate between different nucleotides, even when in complex with their respective R1 partners. Using mutagenesis, we define surfaces on the LlaI.R1 structure that are critical for DNA-binding and interaction with LlaI.R2. These observations support computational modelling of the assembled LlaI restriction system bound to DNA. Together, our data provide new insights into the evolution of nucleotide specificity in McrBC restriction complexes and the molecular mechanisms governing McrBC-catalyzed DNA translocation and cleavage.

14
Accurate and efficient prediction of protein conformations with ProtMonomer

Si, Y.; Zhang, S.; Chen, L.

2026-08-31 molecular biology 10.64898/2026.08.28.747824 medRxiv
Top 1%
0.2%
Show abstract

Deep learning-based protein structure prediction methods that leverage evolutionary information from multiple sequence alignments (MSAs), exemplified by AlphaFold2, have achieved remarkable accuracy. However, existing methods still struggle to predict challenging proteins, particularly those with novel folds or limited evolutionary information, and to recover alternative conformational states. Here we show that structure prediction models trained under different MSA-depth distributions corresponding to different levels of evolutionary information exhibit complementary generalization behaviors, and that a model trained on a mixture of these distributions can combine their complementary generalization strengths. Building on this insight, we developed ProtMonomer, a deep learning framework trained on MSA-depth distributions representing a broad range of evolutionary information levels to improve structure prediction. Across benchmarks comprising CASP15 targets, non-redundant experimentally determined structures, orphan proteins, and short peptides, ProtMonomer performed comparably to or better than leading methods, including AlphaFold2 and AlphaFold3, with particularly strong performance on challenging targets. For fold-switching proteins, ProtMonomer also recovered alternative conformational states more accurately than AlphaFold2 and AlphaFold3 across diverse homologous sequence sampling strategies. In addition to improving predictive accuracy, ProtMonomer substantially reduced inference cost through an efficient architecture, enabling high-throughput applications. Together, these findings provide insights into the generalization of evolution-informed structure prediction models and support ProtMonomer as an accurate and efficient framework for protein structure prediction.

15
The evolutionarily conserved C-terminal domain of a domesticated transposase-derived protein regulates its DNA integration ability

Saha, A.; Ghosh, A.; Majumdar, S.

2026-08-31 biochemistry 10.64898/2026.08.31.747927 medRxiv
Top 2%
0.2%
Show abstract

THAP9 is a transposable element-derived gene which encodes a protein that is homologous to the active Drosophila P-element transposase (DmTNP). Both THAP9 and DmTNP possess a C-terminal domain (CTD) which is functionally uncharacterized. Sequence and structural analysis suggest that the THAP9-CTD has a novel fold which is only found in THAP9 homologs. To explore the evolutionary history and characteristics of this novel domain, exhaustive phylogenetic analysis (using MSA, structure prediction, MSTA-based clustering) was performed. THAP9-CTD homologs were more widely distributed throughout the animal kingdom in comparison to DmTNP-CTD homologs which were restricted to arthropods. Moreover, the THAP9-CTD homologs were more conserved, especially among mammals and birds and their average length increased in a class-specific manner. Comparison with the DmTNP-CTD homologs demonstrates that although their respective CTDs may have evolved independently, they both surprisingly share similar secondary structure elements consisting of three conserved helical regions made of hydrophobic residues that are predicted to make up a conserved core. The role of the respective CTDs were further investigated by creating truncation mutants lacking the CTD. Interestingly both THAP9 and DmTNP truncation mutants are still capable of DNA excision and integration suggesting that their respective CTDs are not essential for DNA transposition. Moreover, CTD truncation favours DNA integration in THAP9: this suggests that CTD acquisition during evolution may have led to THAP9 domestication as observed in other transposable element-derived genes like Rag1 and piggybac, which have similar terminal regulatory domains.

16
Inheritance of a Single Edited CD46 Allele Is Associated with Reduced Ex Vivo Susceptibility to Bovine Viral Diarrhea Virus

Workman, A. M.; Krueger, A. C.; Heaton, M. P.; Snider, A. P.; Kuhn, K. L.; Sonstegard, T. S.; Vander Ley, B. L.

2026-09-01 molecular biology 10.64898/2026.08.31.748238 medRxiv
Top 2%
0.2%
Show abstract

Bovine viral diarrhea virus (BVDV) remains an economically important pathogen of cattle despite widespread vaccination. A homozygous CD46-edited Gir heifer (Ginger) was previously shown to have significantly reduced susceptibility to BVDV. The edited allele contains an in-frame six amino acid substitution within the virus-binding domain of the BVDV entry receptor CD46, replacing residues G82QVLAL with A82LPTFS. Here, we investigated whether reduced BVDV susceptibility is maintained when the edited allele is inherited in the heterozygous state. Ginger was artificially inseminated with semen from an unedited Gir bull and produced a healthy heterozygous CD46-edited bull calf (Giraldo). Whole-genome sequencing confirmed the inheritance and structural integrity of Giraldo's edited allele. Compared with Ginger, Giraldo exhibited similarly reduced ex vivo BVDV susceptibility across primary fibroblasts, lymphocytes, and monocytes, despite inheriting a wild-type CD46 allele from the sire. Allele-specific CD46 RNA expression analysis demonstrated expression of both the edited and wild-type CD46 alleles. Thus, the reduced-susceptibility phenotype was not attributable to transcriptional silencing of the wild-type allele. Lentiviral complementation studies in CD46-knockout Madin-Darby bovine kidney (MDBK) cells further demonstrated that this wild-type CD46 allele was competent to support BVDV infection when expressed independently. Together, these findings indicate that the CD46 A82LPTFS allele can confer reduced BVDV susceptibility in the heterozygous state despite expression of a functional wild-type CD46 allele. This result suggests the potential to more rapidly disseminate reduced BVDV susceptibility through conventional breeding using homozygous CD46-edited sires.

17
FlexiTAC enables controllable PROTAC linker generation across diverse structural settings using a Bayesian flow network with posterior guidance

Li, Y.; Zhao, Y.; Zhou, L.; Huang, C.; Xu, Q.; Chen, Y.; Qin, Z.; Fan, K.; Yang, J.; Cao, D.

2026-08-30 bioinformatics 10.64898/2026.08.26.747172 medRxiv
Top 2%
0.1%
Show abstract

Linker chemistry and conformation are central determinants of PROTAC activity, shaping ternary-complex geometry, cooperativity, target-lysine presentation and cellular permeability. Existing linker generators often lack explicit control over linker flexibility, require predefined attachment sites and linker lengths, or produce structures that demand substantial geometric correction, limiting their utility in practical PROTAC design. Here we introduce FlexiTAC, a Bayesian flow network that jointly generates linker atom types and coordinates from the warhead and E3-ligase-ligand contexts. We also assemble PROTAC-3D, a quality-controlled collection of 63,554 component-resolved PROTAC structures for model training, and PROTAC-Bench, which covers molecular quality, fragment preservation, geometric fidelity, conformational stability, fragment awareness, rediscovery and sampling efficiency. Compared to the best 3D baseline models, FlexiTAC improves validity by 12.0-12.7% and achieves the highest PoseBusters pass rate of 79.5%-80.0%. A differentiable guidance module shifted generated linkers along a conformational ensemble-derived rigidity axis without retraining the generator. In silico case studies further show that the model can accept crystal-derived, redocked or predicted structural inputs. Together, FlexiTAC, PROTAC-3D and PROTAC-Bench establish an integrated and reproducible framework for data-driven PROTAC linker design, combining controllable structure-conditioned generation with standardized training data and evaluation protocols. This framework expands the linker chemical and conformational space accessible to computational exploration, provides a foundation for future method development and enables the systematic generation of structure-conditioned linker designs with tunable conformational flexibility.

18
Structure of an RNA polymerase ribozyme replication complex

Strutzenberg, T. S.; Horning, D. P.; Cochrane, W. G.; Andrade, L.; Han, X.; Joyce, G. F.; Lyumkis, D.

2026-08-31 biophysics 10.64898/2026.08.30.748161 medRxiv
Top 2%
0.1%
Show abstract

Life began with the emergence of a molecule that could replicate its own genetic material, a task plausibly mediated by an RNA-dependent RNA polymerase ribozyme. Here, we present the structure of such a polymerase ribozyme, bound to RNA substrates comprising the template, primer, and nucleoside triphosphate (NTP) analog. The structure reveals how directed evolution shaped flanking elements around a highly conserved catalytic core derived from the ancestral class I ligase ribozyme. Each element serves as a functional module, positioning the primer-template duplex and incoming NTP within the active site of the enzyme. This emergent domain organization is remarkably similar to the "right hand" configuration of polymerase proteins, suggesting a common functional form for copying nucleic acids, regardless of biopolymer catalyst.

19
Rclade: automated taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R

Zeng, Z.; Wang, Y.

2026-09-01 bioinformatics 10.64898/2026.08.27.747462 medRxiv
Top 2%
0.1%
Show abstract

Background: Reproducible taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R often require coordination among several packages and repeated code for label parsing, clade validation, plotting, and export. Workflow-managed analyses additionally benefit from non-interactive configuration, predictable diagnostics, and machine-readable exit status. Results: We present Rclade, an R package that consolidates the multi-package coordination required for taxonomic collapsing into a streamlined, single-function interface. Rclade provides (1) custom ggproto objects (GeomPolygonStraight/GeomSegmentStraight) that bypass coord_munch() interpolation to achieve straight-edge rendering of collapsed triangles in circular layouts; (2) automatic detection and parsing of four taxonomic-label formats (GTDB, Silva, NCBI, embedded) plus user-supplied custom regex, with explicit input-validation contracts and parsing-accuracy evaluation on real and derived test sets; and (3) workflow embeddability through YAML configuration, library-mode APIs, and standard Unix exit codes. Benchmarks on synthetic and real datasets (200-10,000 synthetic tips and real reference trees up to 10,122 tips; 5 replicates at every scale under a unified fully rendered measurement protocol) show that the full-pipeline overhead is modest for interactive use (median {approx}0.87 s in-session rendering and {approx}8.4 s process-level wall-clock at 10,000 tips). Conclusions: Rclade is a convenience layer over the ggtree/deeptime ecosystem that reduces boilerplate while adding targeted technical improvements for circular-layout rendering and format heterogeneity management.

20
PhageTransformer - scalable and accurate host assignments for bacteriophages

Siemers, M.; Lopez, J. L.; Dutilh, B. E.

2026-08-30 bioinformatics 10.64898/2026.08.29.748026 medRxiv
Top 2%
0.1%
Show abstract

Bacteriophages can only be understood through their interactions with bacterial hosts. As environmental sequencing efforts expanded, the number of available phage genome sequences has exploded, yet the vast majority of these sequences lack host information. Predicting the host of a newly observed phage is therefore a key challenge in virology. Several computational tools can predict phage-host relationships from genomic data, but they share notable limitations: (1) the number of different hosts that can be predicted remains relatively restricted; (2) tools tend to assign confident host predictions to non-viral input sequences; and (3) most tools have a trade-off between accuracy and speed. Here we present PhageTransformer (PT), a deep learning model for phage-host prediction that addresses these limitations. We benchmark PT against existing tools on 3,881 independent phage-host pairs from GenBank and public HiC data, and demonstrate that it achieves competitive or superior prediction accuracy at greatly reduced runtime.